Back

Kidney International Reports

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match Kidney International Reports's content profile, based on 15 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Cross-System Meta-Analysis of Machine Learning Predictors Identifies Value-Specific Risk Drivers and Interactions Underlying Acute Kidney Injury

Chan, H. Y.; Li, D.; Yu, A. S. L.; Kellum, J. A.; Fuhrman, D. Y.; Xu, Q.; Chrischilles, E. A.; Cowell, L. G.; Chandaka, S.; Anzalone, A. J.; Kean, J.; McTigue, K. M.; Mosa, A. S. M.; Taylor, B.; Syed, M.; Waitman, L. R.; Hu, Y.; Liu, M.

2026-09-02 nephrology 10.64898/2026.08.31.26361849 medRxiv
Top 0.1%
15.2%
Show abstract

Background: Current understanding of acute kidney injury (AKI) risk factors remains largely descriptive, offering limited precision into how specific biomarker values or physiologic thresholds influence susceptibility. We aimed to synthesize knowledge from machine learning models trained across multiple health systems to identify generalizable, value-specific risk drivers and biomarker interactions contributing to AKI risk. Methods: We analyzed electronic health records (EHRs) from 785,497 adult inpatients between 2010 and 2019 across nine U.S. academic medical centers within PCORnet. Interpretable gradient boosting machine models were independently developed at each health system to quantify predictor-outcome associations. Meta-regression was applied to integrate these site-level results, characterize nonlinear value-risk relationships, and identify bivariate interactions between predictors. Results: Meta-analysis revealed consistent, value-specific risk drivers across health systems. An increase in glucose from 100 mg/dL to 140 mg/dL was associated with a 1.46-fold higher risk of AKI. Chloride and anion gap also demonstrated elevated AKI risk with risk increases overlapping portions of their reference ranges, with anion gap showing a 1.14-fold increase across 4-12 mmol/L and chloride a 1.28-fold increase across 96-100 mEq/L. Electrolytes including potassium, calcium, and sodium showed quadratic associations with AKI risk. Bivariate meta-regression identified interactions between key predictors, highlighting pathways that jointly modulate AKI risk. Conclusion: This cross-system meta-analysis synthesizes machine learning-derived evidence into clinically interpretable knowledge, revealing how specific biomarker ranges and interactions modulate AKI risk. By moving beyond surface-level associations to quantitative, generalizable physiologic thresholds, these findings provide actionable insights to enhance risk stratification and personalized prevention in hospital care.

2
Individual-Level Counterfactual Analysis of SGLT2 Inhibitors Versus DPP4 Inhibitors in Diabetic Kidney Disease Using Causal Machine Learning

Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.

2026-09-03 health informatics 10.64898/2026.08.30.26361750 medRxiv
Top 0.1%
8.0%
Show abstract

Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.

3
Toward Transportable Acute Kidney Injury Prediction: An Explainable XGBoost Model with Temporal Validation Using MIMIC-IV

Okundaye, D. O.; Isiekwene, C. C.

2026-09-03 health informatics 10.64898/2026.09.01.26360393 medRxiv
Top 0.1%
5.4%
Show abstract

Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.

4
Urinary collagen type I degradation products as common fibrosis biomarkers in chronic diseases

Mina, I. K.; Hussain, Y.; Siwy, J.; Catanese, L.; Rupprecht, H.; Beige, J.; Staessen, J. A.; Metzger, J.; Persson, F.; Rossing, P.; Delles, C.; Schanstra, J. P.; Bannaga, A.; Vlahou, A.; Mischak, H.; Arasaradnam, R. P.; Latosinska, A.

2026-08-31 nephrology 10.64898/2026.08.26.26361420 medRxiv
Top 0.1%
5.3%
Show abstract

Background: Fibrosis, characterised by excessive accumulation of collagen type I (COL1), is a common feature of chronic diseases, including liver diseases (LDs), chronic kidney disease (CKD) and heart failure (HF). COL1 degradation products can be detected in urine by proteomics/ peptidomics analyses and may serve as non-invasive biomarkers of fibrosis. We aimed to identify a common molecular signature of fibrosis across these diseases that may ultimately guide interventions to slow disease progression and prevent organ damage. Methods: Using capillary electrophoresis coupled to mass spectrometry (CE-MS), naturally occurring COL1 degradation products (peptides) in the urine of patients with fibrotic disease, LDs (n=127), CKD (n=263) or HF (n=187), were investigated and compared with the same number of matched controls. Disease-associated COL1 peptides were identified separately for each condition, and peptides showing consistent associations across the three diseases were selected to define a common fibrosis signature. A support vector machine model based on the selected peptides was developed and validated in independent cohorts of patients with LDs (n=110), CKD (n=93), HF (n=32) and controls (n=643). Results: We identified a common fibrotic signature consisting of 50 COL1 degradation products, mainly downregulated in fibrosis. A model based on these peptides achieved a strong performance, with an area under the receiver operating characteristic curve (AUC) of 0.935 (95% confidence interval (CI) 0.917-0.953, p<0.0001) in an external validation cohort comprising pooled disease groups (LDs, CKD, and HF) and controls. Performance was maintained in LDs, CKD and HF, with AUCs of 0.917 (95% CI 0.890-0.944, p<0.0001), 0.951 (95% CI 0.931-0.971, p<0.0001) and 0.950 (95% CI 0.903-0.997, p<0.0001), respectively. The model scores were significantly associated with fibrosis stage in LDs (p=0.0097) and with interstitial fibrosis and tubular atrophy in CKD (p=0.045). Conclusion: A model of urinary COL1 peptides captures a shared collagen degradation signature across organs and diseases, enabling the non-invasive assessment of fibrosis irrespective of its origin. As these peptides exclusively reflect collagen degradation, the findings suggest impaired collagen degradation as a driver in fibrosis. Future clinical studies are warranted to evaluate the utility of this model for early fibrosis detection and earlier implementation of anti-fibrotic interventions.

5
BanffNET, a Deep Learning System for Comprehensive Histological Lesion Quantification in Kidney Transplant Biopsies

Buzzanca, G.; Pala, C.; He, J.; Hofstraat-Boersma, R.; Tammaro, A.; van Midden, D.; Buelow, R.; Hoelscher, D. L.; Muehlfeld, A. S.; Koeller, m.; Kozakowski, N.; Boehmig, G.; Halloran, P. F.; van der Helm, D.; Meziyerh, S.; Venhuizen, J.-H.; Haitjema, S.; Dijkstra, J.; Hilbrands, L. B.; Steenbergen, E. J.; van Zuilen, A. D.; Nurmohamed, A. S.; Bemelman, F. J.; Bruns, I. B.; Callegaro, G.; van de Water, B.; Pieters, T. T.; Breimer, G. E.; Rossi, G. M.; Fiaccadori, E.; Maggiore, U.; Roelofs, J. J. T. H.; Testa, F.; Fontana, F.; Abiola, A. A.; Delsante, M.; Corthals, G. L.; Peters-Sengers, H.; Ngu

2026-09-02 pathology 10.64898/2026.08.28.26360029 medRxiv
Top 0.1%
4.9%
Show abstract

Accurate, reproducible interpretation of kidney allograft biopsies is critical for diagnosis of graft injury to guide prognosis and management. The international Banff classification is a consensus diagnostic system based on semiquantitative histological lesion scoring on either extent or severity of kidney transplant biopsies. However, pathologist scoring is limited by substantial interobserver variability, constrained scalability, and the inherent nature of the scoring system itself. Here we present BanffNET, a weakly supervised, probabilistic deep learning framework that combines self-supervised feature extraction with a novel Bayesian multiple-instance learning framework to predict (continuously) the full spectrum of Banff lesion scores directly from whole-slide images (WSIs). Using lesion-specific aggregation functions tailored to localized (modeling lesion severity) and diffuse pathologies (modeling lesion extent), BanffNET generates interpretable, patch-level probability maps and calibrated slide-level scores. BanffNET's performance was assessed relative to consensus, biological correlates of rejection and clinical outcome, demonstrating superior consistency, transportability and generalization. Trained on 7,249 WSIs from three cohorts, BanffNET demonstrates consistent performance on 11,028 WSIs across five external test sets, performing on par or exceeding expert consensus across lesions. BanffNET scores align more closely than pathologist Banff scores with molecular profiles of rejection, offering a transparent, biologically grounded framework for computational pathology with relevance beyond transplantation.

6
Belimumab with rituximab for the treatment of primary membranous nephropathy

Chung, S. A.; Stelzig, L.; Sherman, M. A.; Gao, W.; Tosta, P.; Cooney, L. A.; Adler, S.; Aslam, N.; Ayoub, I.; Bomback, A. S.; Coppock, G.; Derebail, V. K.; Kamal, F.; Rizk, D. V.; Tuttle, K. R.; Waldman, M.; Barry, W. T.; Nachman, P. H.

2026-08-31 nephrology 10.64898/2026.08.26.26360913 medRxiv
Top 0.1%
4.4%
Show abstract

Introduction: B cell depletion with rituximab leads to complete or partial remission (CR/PR) in only ~60% of patients with primary membranous nephropathy (PMN). Adding belimumab to rituximab may result in greater depletion of memory B cells, limit the re-emergence of autoreactive B cells, and improve clinical responses. Methods: REBOOT Part A (NCT03949855) is a single arm, open-label, pharmacokinetic study where all participants had proteinuria [&ge;] 4g/day and detectable serum anti-phospholipase A2 receptor (anti-PLA2R) antibodies. Participants received belimumab 200 mg subcutaneously weekly for 52 weeks and rituximab 1000 mg intravenously at weeks 4 and 6. Assessments included belimumab exposure at week 4 and CR/PR at week 104. Results: Seventeen participants started belimumab. Belimumab exposure was not significantly reduced in those with high (> 9 g/day) proteinuria at week 4. Among all treated participants, 59% (10/17) achieved CR/PR at week 104, while in per protocol analyses, 91% (10/11) achieved CR/PR at week 104. All participants in per protocol analyses had normal serum albumin and undetectable serum anti-PLA2R by week 104. Circulating memory B cells increased before rituximab and were depleted by rituximab. B cell re-constitution occurred after week 52 with primarily naive and transitional B cells. Belimumab with rituximab was well-tolerated, with one participant discontinuing belimumab due to infection. Conclusion: In this study, a high proportion of participants receiving belimumab with rituximab achieved CR/PR. Thus, a multi-targeted approach to B cell depletion may improve immunologic and clinical outcomes in PMN and is being studied in a larger, randomized, placebo-controlled clinical trial.

7
Sex Differences in the Impact of Allosensitization on Waitlist Access and Post-Transplant Outcomes in Adults with Congenital Heart Disease

Joseph, A.; Kearney, K.; Henricks, C.; Morgan, J. L.; Tan, W.; Shafer, K.; Wrobel, C.; Lacelle, C.; Burns, K.; Jawaid, A.; Tapaskar, N.; Solmonson, A.; Nelson, D. B.; Truby, L. K.

2026-09-02 transplantation 10.64898/2026.08.31.26361832 medRxiv
Top 0.2%
1.4%
Show abstract

Background: Adult congenital heart disease (ACHD) patients are prone to HLA-antibody formation from multiple surgeries, transfusions, and prosthetic surgical material. Females with ACHD may accrue additional, non-surgical alloantigen exposure. Whether sex modifies the impact of allosensitization on heart transplant (HT) access and outcomes in ACHD remains unknown. Methods: We retrospectively analyzed the OPTN/UNOS registry of adults with ACHD listed for first-time HT (2018-2025). Sensitization was defined by calculated panel reactive antibodies (cPRA) at listing. We tested the sex x sensitization (highly sensitized, cPRA >50%) interaction on transplant access using Fine-Gray competing-risks regression, treating transplantation as the event of interest and death or removal from the waitlist as competing events, and on post-transplant survival using multivariable Cox proportional-hazards regression, both adjusted for age at listing, mechanical support at listing, and the number of distinct prior cardiac surgery categories. Results: Among 856 candidates (38% female), females were more often highly sensitized than males (23% vs 14%; age-adjusted OR 1.81, 95% CI 1.26-2.61), even after adjusting for surgical burden. Sensitization reduced transplant access in females (84% to 71%; median wait 60 to 110 days, p < 0.001) but not males (79% vs 79%, median wait 88 vs 98 days). In adjusted Fine-Gray models, the subdistribution hazard for transplant was reduced in sensitized females (sHR 0.54, 95% CI 0.41-0.72) with no effect in males (sHR 0.96, 95% CI 0.73-1.26), and the sex x sensitization interaction was significant (interaction sHR 0.64, 95% CI 0.44-0.94, p = 0.02). Post-transplant mortality was numerically higher in sensitized than non-sensitized candidates in both sexes and the sex x sensitization interaction on 1-year mortality was not significant. The sex-asymmetric effect persisted and was more pronounced in the multiorgan candidates. Conclusions: Allosensitization is not a sex-neutral barrier to transplant in HT candidates with ACHD. Females are more sensitized and have reduced transplant access without differences in 1-year mortality. The female excess in sensitization is not accounted for by surgical burden, and the exposures responsible remain to be defined. These findings warrant a sex-aware listing strategy and further studies.

8
Validation of individualized flow simulations for determining the pressure gradient in patients with renal artery stenosis

Bouwmeester, T. A.; Collard, D.; Zijlstra, I. A. J.; van Hulst, E.; Lamers, A. G. B. H.; Vogt, L.; van den Born, B.-J. H.; van de Velde, L.

2026-08-31 radiology and imaging 10.64898/2026.08.27.26361537 medRxiv
Top 0.2%
1.2%
Show abstract

Objectives To validate two computational fluid dynamics (CFD) models derived from computed tomography angiography (CTA) for estimating trans-stenotic pressure gradients, using invasive intra-arterial pressure measurements as the reference standard in patients with renal artery stenosis (RAS). Background We assessed whether non-invasive assessment of the pressure gradient using CFD could be a reliable alternative to intra-arterial measurements for identifying hemodynamically significant RAS. Methods We performed intra-arterial measurements at rest and during dopamine-induced hyperemia to assess the trans-stenotic pressure gradient in 28 patients with RAS. A pre-intervention CTA scan was used to simulate the pressure gradient with a CFD model using a strategy based on Murray's law (CFD-Mu) and cortical volume (CFD-C). The agreement between the simulated and measured pressure gradients was assessed using intraclass correlation coefficients (ICC), Bland-Altman analysis and diagnostic agreement on the presence of a hemodynamically significant stenosis. Results In 20 patients, successful measurements and simulations were obtained. The ICC between measured pressure gradient and the CFD pressure gradient was 0.78 and 0.94 during baseline and 0.86 and 0.72 during hyperemia, for CFD-Mu and CFD-C, respectively. The sensitivity of CFD-Mu and CFD-C was 70% for both models at rest and 100% compared to the hyperemic measurements, whereas the specificity was 90% and 70% at rest and 79% and 72% during hyperemia, respectively. Conclusions The results support the use of individualized CFD simulations for hemodynamic assessment of RAS using CTA as input. The CFD models demonstrated high accuracy for the identification of a hemodynamically significant stenosis.

9
Long-Term Impact of Cumulative Hyperglycaemia on DNA Methylation and its Role in Diabetic Kidney Disease

Luo, X.; Syreeni, A.; Hill, C.; Smyth, L. J.; Dahlstrom, E. H.; Mutter, S.; Chen, Z.; Natarajan, R.; Pan, S.; Parton, A.; Jackson, H.; McKay, G.; Susztak, K.; Hirschhorn, J. N.; Florez, J. C.; Maxwell, A. P.; Groop, P.-H.; McKnight, A. J.; Sandholm, N.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.31.26361614 medRxiv
Top 0.4%
0.3%
Show abstract

Hyperglycaemia is a hallmark of diabetes and a major risk factor for diabetic kidney disease (DKD). However, the molecular consequences of long-term cumulative hyperglycaemia (CH) remain unclear. As a stable epigenetic modification, DNA methylation may capture past glycaemic exposure. Here, we assessed CH-associated DNA methylation in 1,245 participants with type 1 diabetes (T1D) from Finland and the United Kingdom-Republic of Ireland cohorts. We identified 17 CH-associated CpGs, with the strongest association at cg19693031 (TXNIP). Longitudinal analyses demonstrate that these CH-associated DNA methylation levels remain stable despite short-term glycaemic fluctuations, suggesting lasting epigenetic imprints of earlier metabolic control. Integrative analyses combining genomic, epigenetic, and proteomic data characterized these CpGs and potential target proteins. Mendelian randomization suggested a causal association between cg20853880 (KLF11) and DKD, supported by chromatin accessibility and kidney KLF11 expression. Our findings suggest that epigenetic changes contribute to metabolic memory and may mediate the effects of hyperglycaemia on DKD.

10
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.4%
0.3%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

11
Predicting COVID-19 hospitalisation and common disease risk from comorbid diagnoses in 13 million individuals

Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,

2026-09-01 health informatics 10.64898/2026.08.27.26361302 medRxiv
Top 0.5%
0.2%
Show abstract

Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.

12
AURORA: Analysing and understanding responses to oncological regimens with artificial intelligence

Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.

2026-09-02 health informatics 10.64898/2026.08.30.26361778 medRxiv
Top 0.5%
0.2%
Show abstract

Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.

13
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 0.6%
0.2%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

14
ECG-based longitudinal risk prediction across diseases and organ systems

ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.

2026-09-02 health informatics 10.64898/2026.08.29.26361697 medRxiv
Top 0.7%
0.1%
Show abstract

Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.

15
A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders

Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.

2026-08-31 health informatics 10.64898/2026.08.28.26361655 medRxiv
Top 0.8%
0.1%
Show abstract

Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.

16
People living with multiple long-term conditions have different pathways of unscheduled care in hospital: findings from an analysis of routinely-collected clinical data

Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.

2026-09-01 health informatics 10.64898/2026.08.28.26361696 medRxiv
Top 0.8%
0.1%
Show abstract

Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.

17
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 1.0%
0.1%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

18
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 1%
0.1%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

19
A Curated Pharmacogenomic Allele Catalog for Sub-Saharan African Populations

SULAIMAN, M. A.; Oyeyemi, B. F.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361354 medRxiv
Top 1%
0.1%
Show abstract

Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.

20
PCGS: biomarker and risk group identification for Pediatric Cancers via explainable Graph neural networks with Shapley values

Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.

2026-09-01 health informatics 10.64898/2026.08.27.26361540 medRxiv
Top 1%
0.1%
Show abstract

Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.